Deploy and manage Weave on your own infrastructure
Weave on self-managed infrastructure is currently in Private Preview.For production environments, W&B strongly recommends using W&B Dedicated Cloud, where Weave is Generally Available.To deploy a production-grade, self-managed instance, contact support@wandb.com.
This guide explains how to deploy all the components required to run W&B Weave in a self-managed environment.A key component of a self-managed Weave deployment is ClickHouseDB, which the Weave application backend relies on.Although the deployment process sets up a fully functional ClickHouseDB instance, you may need to take additional steps to ensure reliability and high availability for a production-ready environment.
The ClickHouse deployment in this document uses the Bitnami ClickHouse package.The Bitnami Helm chart provides good support for basic ClickHouse functionalities, particularly the use of ClickHouse Keeper.To configure Clickhouse, complete the following steps:
The most critical part of the Helm configuration is the ClickHouse configuration, which is provided in XML format. Below is an example values.yaml file with customizable parameters to suit your needs.
To make the configuration process easier, we have added comments in the relevant sections using the format {/* COMMENT */}.Modify the following parameters:
clusterName
auth.username
auth.password
S3 bucket-related configurations
W&B recommends keeping the clusterName value in values.yaml set to weave_cluster. This is the expected cluster name when W&B Weave runs the database migration. If you need to use a different name, see the Setting clusterName section for more information.
## @param clusterName ClickHouse cluster nameclusterName: weave_cluster## @param shards Number of ClickHouse shards to deployshards: 1## @param replicaCount Number of ClickHouse replicas per shard to deploy## if keeper enable, same as keeper count, keeper cluster by shards.replicaCount: 3persistence: enabled: true size: 30G # this size must be larger than cache size.## ClickHouse resource requests and limitsresources: requests: cpu: 0.5 memory: 500Mi limits: cpu: 3.0 memory: 6Gi## Authenticationauth: username: weave_admin password: "weave_123" existingSecret: "" existingSecretKey: ""## @param logLevel Logging levellogLevel: information## @section ClickHouse keeper configuration parameterskeeper: enabled: true## @param extraEnvVars Array with extra environment variables to add to ClickHouse nodes##extraEnvVars: - name: S3_ENDPOINT value: "https://s3.us-east-1.amazonaws.com/bucketname/$(CLICKHOUSE_REPLICA_ID)"## @param defaultConfigurationOverrides [string] Default configuration overrides (evaluated as a template)defaultConfigurationOverrides: | <clickhouse> {/* Macros */} <macros> <shard from_env="CLICKHOUSE_SHARD_ID"></shard> <replica from_env="CLICKHOUSE_REPLICA_ID"></replica> </macros> {/* Log Level */} <logger> <level>{{ .Values.logLevel }}</level> </logger> {{- if or (ne (int .Values.shards) 1) (ne (int .Values.replicaCount) 1)}} <remote_servers> <{{ .Values.clusterName }}> {{- $shards := $.Values.shards | int }} {{- range $shard, $e := until $shards }} <shard> <internal_replication>true</internal_replication> {{- $replicas := $.Values.replicaCount | int }} {{- range $i, $_e := until $replicas }} <replica> <host>{{ printf "%s-shard%d-%d.%s.%s.svc.%s" (include "common.names.fullname" $ ) $shard $i (include "clickhouse.headlessServiceName" $) (include "common.names.namespace" $) $.Values.clusterDomain }}</host> <port>{{ $.Values.service.ports.tcp }}</port> </replica> {{- end }} </shard> {{- end }} </{{ .Values.clusterName }}> </remote_servers> {{- end }} {{- if .Values.keeper.enabled }} <keeper_server> <tcp_port>{{ $.Values.containerPorts.keeper }}</tcp_port> {{- if .Values.tls.enabled }} <tcp_port_secure>{{ $.Values.containerPorts.keeperSecure }}</tcp_port_secure> {{- end }} <server_id from_env="KEEPER_SERVER_ID"></server_id> <log_storage_path>/bitnami/clickhouse/keeper/coordination/log</log_storage_path> <snapshot_storage_path>/bitnami/clickhouse/keeper/coordination/snapshots</snapshot_storage_path> <coordination_settings> <operation_timeout_ms>10000</operation_timeout_ms> <session_timeout_ms>30000</session_timeout_ms> <raft_logs_level>trace</raft_logs_level> </coordination_settings> <raft_configuration> {{- $nodes := .Values.replicaCount | int }} {{- range $node, $e := until $nodes }} <server> <id>{{ $node | int }}</id> <hostname from_env="{{ printf "KEEPER_NODE_%d" $node }}"></hostname> <port>{{ $.Values.service.ports.keeperInter }}</port> </server> {{- end }} </raft_configuration> </keeper_server> {{- end }} {{- if or .Values.keeper.enabled .Values.zookeeper.enabled .Values.externalZookeeper.servers }} <zookeeper> {{- if or .Values.keeper.enabled }} {{- $nodes := .Values.replicaCount | int }} {{- range $node, $e := until $nodes }} <node> <host from_env="{{ printf "KEEPER_NODE_%d" $node }}"></host> <port>{{ $.Values.service.ports.keeper }}</port> </node> {{- end }} {{- else if .Values.zookeeper.enabled }} {{- $nodes := .Values.zookeeper.replicaCount | int }} {{- range $node, $e := until $nodes }} <node> <host from_env="{{ printf "KEEPER_NODE_%d" $node }}"></host> <port>{{ $.Values.zookeeper.service.ports.client }}</port> </node> {{- end }} {{- else if .Values.externalZookeeper.servers }} {{- range $node :=.Values.externalZookeeper.servers }} <node> <host>{{ $node }}</host> <port>{{ $.Values.externalZookeeper.port }}</port> </node> {{- end }} {{- end }} </zookeeper> {{- end }} {{- if .Values.metrics.enabled }} <prometheus> <endpoint>/metrics</endpoint> <port from_env="CLICKHOUSE_METRICS_PORT"></port> <metrics>true</metrics> <events>true</events> <asynchronous_metrics>true</asynchronous_metrics> </prometheus> {{- end }} <listen_host>0.0.0.0</listen_host> <listen_host>::</listen_host> <listen_try>1</listen_try> <storage_configuration> <disks> <s3_disk> <type>s3</type> <endpoint from_env="S3_ENDPOINT"></endpoint> {/* AVOID USE CREDENTIALS CHECK THE RECOMMENDATION */} <access_key_id>xxx</access_key_id> <secret_access_key>xxx</secret_access_key> {/* AVOID USE CREDENTIALS CHECK THE RECOMMENDATION */} <metadata_path>/var/lib/clickhouse/disks/s3_disk/</metadata_path> </s3_disk> <s3_disk_cache> <type>cache</type> <disk>s3_disk</disk> <path>/var/lib/clickhouse/s3_disk_cache/cache/</path> {/* THE CACHE SIZE MUST BE LOWER THAN PERSISTENT VOLUME */} <max_size>20Gi</max_size> </s3_disk_cache> </disks> <policies> <s3_main> <volumes> <main> <disk>s3_disk_cache</disk> </main> </volumes> </s3_main> </policies> </storage_configuration> <merge_tree> <storage_policy>s3_main</storage_policy> </merge_tree> </clickhouse>## @section Zookeeper subchart parameterszookeeper: enabled: false
Do not remove the $(CLICKHOUSE_REPLICA_ID) from the bucket endpoint configuration. It will ensure each ClickHouse replica is writing and reading data from it’s folder in the bucket.
You can specify credentials for accessing an S3 bucket by either hardcoding the configuration, or having ClickHouse fetch the data from environment variables or an EC2 instance.
Replace <release-name> with your Helm release name
Replace <namespace> with your NAMESPACE
Get the service details: kubectl get svc -n <namespace>
Username: Set in the values.yaml
Password: Set in the values.yaml
With this information, update the W&B Platform Custom Resource(CR) by adding the following configuration:
apiVersion: apps.wandb.com/v1kind: WeightsAndBiasesmetadata: labels: app.kubernetes.io/name: weightsandbiases app.kubernetes.io/instance: wandb name: wandb namespace: defaultspec: values: global: [...] clickhouse: host: <release-name>-headless.<namespace>.svc.cluster.local port: 8123 password: <password> user: <username> database: wandb_weave # `replicated` must be set to `true` if replicating data across multiple nodes # This is in preview, use the env var `WF_CLICKHOUSE_REPLICATED` replicated: true weave-trace: enabled: true [...] weave-trace: install: true extraEnv: WF_CLICKHOUSE_REPLICATED: "true" [...]
When using more than one replica (W&B recommend a least 3 replicas), ensure to have the following environment variable set for Weave Traces.
extraEnv: WF_CLICKHOUSE_REPLICATED: "true"
This has the same effect of replicated: true which in preview.
Set the clusterName in values.yaml to weave_cluster. If it is not, the database migration will fail.Alternatively, ff you use a different cluster name, set the WF_CLICKHOUSE_REPLICATED_CLUSTER environment variable in weave-trace.extraEnv to match the chosen name, as shown in the example below.
[...] clickhouse: host: <release-name>-headless.<namespace>.svc.cluster.local port: 8123 password: <password> user: <username> database: wandb_weave # `replicated` must be set to `true` if replicating data across multiple nodes # This is in preview, use the env var `WF_CLICKHOUSE_REPLICATED` replicated: true weave-trace: enabled: true[...]weave-trace: install: true extraEnv: WF_CLICKHOUSE_REPLICATED: "true" WF_CLICKHOUSE_REPLICATED_CLUSTER: "different_cluster_name"[...]
The final configuration will look like the following example: